Back

Annals of Internal Medicine

American College of Physicians

Preprints posted in the last 90 days, ranked by how well they match Annals of Internal Medicine's content profile, based on 28 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.

1
Herpes Simplex Virus Serostatus and 1-Year Mortality Among Allogeneic Hematopoietic Cell Transplant Recipients

Fischer, M. D.; Johnston, C.; Boeckh, M. J.; Ford, E. S.; Gooley, T.; Phipps, A. I.; Winer, R. L.; Biernacki, M. A.; McCulloch, D. J.; Sandmaier, B. M.; Greninger, A. L.; Wald, A.; Pergam, S. A.

2026-08-02 transplantation 10.64898/2026.07.30.26359369 medRxiv
Top 0.1%
12.0%
Show abstract

Viral infections remain a cause of substantial morbidity and mortality in allogeneic hematopoietic cell transplant (aHCT) recipients. While antiviral prophylaxis has dramatically reduced the risk of herpes simplex virus (HSV) disease, the relationship between HSV serostatus and major post-transplant complications in the context of HSV prophylaxis is unknown. We evaluated the association between HSV serostatus and survival among adults who received a first aHCT at the Fred Hutchinson Cancer Center between 2002 and 2022. Patients were screened for HSV-1 and HSV-2 by Western blot (WB) prior to transplant. We fit Cox proportional hazards models for mortality up to one year post-transplant, comparing HSV seropositive to seronegative patients. Models were adjusted for age, sex, cytomegalovirus (CMV) serostatus, conditioning regimen, disease risk, graft type and HLA matching, year of transplant and acute graft-versus host disease. A total of 4,016 aHCT recipients were included in this analysis. The cumulative all-cause 1-year mortality was 29.8%. For HSV-1, the adjusted hazard ratio (aHR) for all-cause mortality comparing seropositive to seronegative individuals was 1.20 (95% CI: 1.06-1.36). The aHRs for relapse, non-relapse mortality (NRM) and relapse-related mortality (RRM) were 1.46 (1.24-1.71), 1.04 (0.89-1.21), and 1.75 (1.40-2.19), respectively. For HSV-2, the aHRs for all-cause mortality, relapse, NRM, and RRM were 1.03 (0.93-1.14), 1.16 (1.03-1.31), 0.99 (0.87-1.13), and 1.13 (0.97-1.33), respectively. Despite universal antiviral prophylaxis, HSV-1 seropositivity was associated with higher mortality in the year after transplant, driven by RRM. Further studies are needed to confirm the association and understand the potential mechanisms underlying this relationship.

2
Clinical characteristics and associated factors of de novo and recurrent prostate cancer after kidney transplantation

Apanisile, K.; Li, M.-H.; Faddoul, G.; Ekwenna, O.; Koizumi, N.

2026-08-02 transplantation 10.64898/2026.07.30.26359389 medRxiv
Top 0.1%
9.6%
Show abstract

Kidney transplant recipients experience a higher burden of several malignancies, yet the factors associated with prostate cancer presentation after transplantation remain poorly understood. Unlike malignancies strongly associated with impaired immune surveillance, prostate cancer has not consistently demonstrated an increased incidence after transplantation, suggesting that different mechanisms may underlie disease presentation. This study evaluated recipient, donor, transplant, immunologic, and immunosuppressive factors associated with prostate cancer phenotype after kidney transplantation. A retrospective cohort study was conducted using national transplant registry data from adult kidney transplant recipients diagnosed with post-transplant prostate cancer between 2015 and 2024. Cases were classified as de novo (no pre-transplant history of prostate cancer) or recurrent (documented pre-transplant history). Multivariable Firth penalized logistic regression was used to evaluate factors associated with recurrent phenotype. Prespecified sensitivity analyses included deceased donor restricted models, incorporation of donor organ quality variables, and adjustment for time from transplantation to cancer diagnosis. Exploratory machine learning analyses included elastic net logistic regression, random forest, and extreme gradient boosting. The cohort included 660 recipients, of whom 623 (94.4%) had de novo disease and 37 (5.6%) had recurrent disease. Recipient age was the only variable consistently associated with recurrent phenotype across primary and sensitivity analyses (adjusted odds ratio per year 1.11, 95% CI 1.05-1.17; p<0.001). Immunosuppressive regimen, donor characteristics, immunologic variables, and time from transplantation to cancer diagnosis were not independently associated with phenotype in the primary cohort. In deceased donor restricted analyses, alemtuzumab induction showed an exploratory association with recurrent phenotype, although estimates were imprecise. Machine learning models demonstrated modest discrimination and calibration and did not outperform penalized regression approaches. These findings suggest that, among kidney transplant recipients with prostate cancer, differences between recurrent and de novo presentation are more closely associated with recipient age and underlying disease characteristics than with transplant exposures or specific immunosuppressive regimens.

3
Trends in NIH funding for pediatric research, FY2020-FY2026

Ledley, F. D.; Mozer, R.

2026-07-01 pediatrics 10.64898/2026.06.29.26356863 medRxiv
Top 0.1%
8.0%
Show abstract

There were substantial changes in NIH policies regarding research funding in FY2025. This work examines NIH funding for pediatric research FY2020 through Q2FY2026 including the number and cost of awards, the number of first year (type 1) awards, the number of Notices of Funding Opportunity, and the topical focus of research awards. NIH funding for pediatric research declined >20% in the first two quarters FY2025-FY2026 with proportionally greater reductions in first year awards and Notices of Funding Opportunity. Changes were also noted in the topic prevalence of new awards consistent with 2025 guidelines identifying topics "not aligned with NIH priorities." These results suggest that pediatric research aimed at advancing healthcare for children is at risk with potential collateral consequences beyond childhood.

4
Determinants, strategies, and outcomes of implementing an enhanced dosimetry quality assurance checklist in radiation oncology: A qualitative implementation science study

Adapa, K.; Mosaly, P. R.; Yu, F.; Moore, C.; McGurk, R.; Das, S.; Mazur, L.

2026-08-10 health informatics 10.64898/2026.08.05.26359837 medRxiv
Top 0.1%
7.6%
Show abstract

Radiation oncology has a long history of developing in-house health information technology (HIT) tools such as quality assurance (QA) checklists, yet there is little guidance from professional bodies on how to implement these tools in complex clinical environments. Building on our previous work that used human-centered participatory co-design, the Task-User-Representation-Function (TURF) framework, and multi-method usability evaluations to design and develop an enhanced dosimetry QA checklist (DQC), this study investigated the barriers and facilitators (determinants) to implementing the enhanced DQC in a radiation oncology clinic, examined implementation strategies, proposed an implementation framework for QA checklists in radiation oncology, and assessed four implementation outcomes: acceptability, appropriateness, feasibility, and adoption. We conducted a qualitative implementation study using an abductive research approach at an academic medical center. All key stakeholders (dosimetrists, physicists, trainees, and software developers) participated in semi-structured interviews, field observations, and surveys across pre-implementation, implementation, and post-implementation phases. Data were analyzed using a hybrid inductive-deductive approach, with deductive coding guided by an adapted Consolidated Framework for Implementation Research (CFIR) mapped to the Unified Theory of Acceptance and Use of Technology and by the Expert Recommendations for Implementing Change (ERIC) compilation. We identified 4 CFIR constructs and 12 sub-constructs as barriers, with structural characteristics and planning showing the highest negative valence, and 5 CFIR constructs and 19 sub-constructs as facilitators, with relative advantage, culture, and leadership engagement showing the highest positive valence. Participants' suggestions mapped to 19 ERIC strategies in 7 clusters, and the CFIR-ERIC matching tool identified 14 evidence-based strategies in 4 clusters that informed a proposed phased implementation framework. Acceptability, appropriateness, and feasibility scores improved significantly from pre-implementation to implementation for all professional roles (p<0.05), yet adoption reached 100% only in the sixth week of implementation. These findings highlight the value of combining subjective and objective implementation outcomes and provide a practical, evidence-based framework for implementing in-house QA checklists in radiation oncology that warrants validation in diverse settings.

5
Combining VEGFR tyrosine kinase inhibitors and PD-1/PD-L1 inhibitors versus VEGFR tyrosine kinase inhibitors monotherapy in renal cell carcinoma: a target trial emulation

Shi, D.; Li, X.; Chen, Y.; Chen, Y.; Song, Q.; Su, J.

2026-07-02 health informatics 10.64898/2026.06.30.26356941 medRxiv
Top 0.1%
7.2%
Show abstract

Importance: Combinations of VEGFR tyrosine kinase inhibitors (TKIs) and immune checkpoint inhibitors (ICIs), such as antibodies to programmed cell death-1 (PD-1), or to its ligand PD-L1, are now first-line standard of care for renal cell carcinoma (RCC), but the pivotal clinical trials excluded patients with common comorbidities, leaving their real-world effectiveness uncertain. Objective: To determine whether adding PD-1/PD-L1 inhibitors to VEGFR-TKIs therapy is associated with improved overall survival in a real-world RCC cohort. Design, Setting, and Participants: This retrospective cohort study used a target trial emulation framework and real-world electronic health records data from the University of Florida Health Integrated Data Repository (IDR). Data was analyzed from September 2009 through June 2023. Adult patients ([&ge;]18 years) with confirmed RCC and at least one VEGFR-TKIs prescription were eligible. The date of the first VEGFR-TKIs prescription was defined as the index date, and patients were followed for up to 24 months. Variable-ratio propensity score matching (up to 2:1) across 13 baseline covariates was used to emulate randomized treatment assignments. Of 107,783 patients screened, 387 met eligibility criteria, and 319 remained in the matched cohort. Exposures: VEGFR-TKIs monotherapy (control group) versus VEGFR-TKIs combined with PD-1/PD-L1 inhibitors (experimental group). Main Outcomes and Measures: Overall survival (OS), analyzed by weighted Kaplan-Meier estimation, cluster-robust Cox regression, and restricted mean survival time (RMST) at {tau} = 24 months, prespecified given anticipated non-proportional hazards. Results: Among 319 matched patients (mean [SD] age, 62 [12] years; 76% male), 107 deaths occurred (33.5%). Twelve-month OS was higher in the combination arm (81.8%; 95% CI, 74.7--89.6%) than VEGFR-TKIs monotherapy (68.1%; 95% CI, 61.1--76.0%), converging by 24 months (61.1% vs 56.7%). The Cox hazard ratio was 0.718 (95% CI, 0.484-- 1.064; P = 0.0986). RMST was 2.79 months greater with combination therapy (95% CI, 0.93-- 4.65; P = 0.0033). Conclusions: Adding PD-1/PD-L1 inhibitors to VEGFR-TKIs therapy was associated with a statistically significant and clinically meaningful gain in restricted mean survival, supporting the real-world generalizability of combination therapy and the importance of appropriate treatment effect measures under non-proportional hazards.

6
Sex Differences in the Impact of Allosensitization on Waitlist Access and Post-Transplant Outcomes in Adults with Congenital Heart Disease

Joseph, A.; Kearney, K.; Henricks, C.; Morgan, J. L.; Tan, W.; Shafer, K.; Wrobel, C.; Lacelle, C.; Burns, K.; Jawaid, A.; Tapaskar, N.; Solmonson, A.; Nelson, D. B.; Truby, L. K.

2026-09-02 transplantation 10.64898/2026.08.31.26361832 medRxiv
Top 0.1%
6.3%
Show abstract

Background: Adult congenital heart disease (ACHD) patients are prone to HLA-antibody formation from multiple surgeries, transfusions, and prosthetic surgical material. Females with ACHD may accrue additional, non-surgical alloantigen exposure. Whether sex modifies the impact of allosensitization on heart transplant (HT) access and outcomes in ACHD remains unknown. Methods: We retrospectively analyzed the OPTN/UNOS registry of adults with ACHD listed for first-time HT (2018-2025). Sensitization was defined by calculated panel reactive antibodies (cPRA) at listing. We tested the sex x sensitization (highly sensitized, cPRA >50%) interaction on transplant access using Fine-Gray competing-risks regression, treating transplantation as the event of interest and death or removal from the waitlist as competing events, and on post-transplant survival using multivariable Cox proportional-hazards regression, both adjusted for age at listing, mechanical support at listing, and the number of distinct prior cardiac surgery categories. Results: Among 856 candidates (38% female), females were more often highly sensitized than males (23% vs 14%; age-adjusted OR 1.81, 95% CI 1.26-2.61), even after adjusting for surgical burden. Sensitization reduced transplant access in females (84% to 71%; median wait 60 to 110 days, p < 0.001) but not males (79% vs 79%, median wait 88 vs 98 days). In adjusted Fine-Gray models, the subdistribution hazard for transplant was reduced in sensitized females (sHR 0.54, 95% CI 0.41-0.72) with no effect in males (sHR 0.96, 95% CI 0.73-1.26), and the sex x sensitization interaction was significant (interaction sHR 0.64, 95% CI 0.44-0.94, p = 0.02). Post-transplant mortality was numerically higher in sensitized than non-sensitized candidates in both sexes and the sex x sensitization interaction on 1-year mortality was not significant. The sex-asymmetric effect persisted and was more pronounced in the multiorgan candidates. Conclusions: Allosensitization is not a sex-neutral barrier to transplant in HT candidates with ACHD. Females are more sensitized and have reduced transplant access without differences in 1-year mortality. The female excess in sensitization is not accounted for by surgical burden, and the exposures responsible remain to be defined. These findings warrant a sex-aware listing strategy and further studies.

7
Algorithmic Ascertainment of Cause of Death from Longitudinal Real-World Medical Claims Data: Development and Validation

McLean, K. W.; LaBonte, J.; Macaulay, K.; Kassam-Adams, S.

2026-08-21 health informatics 10.64898/2026.08.18.26360606 medRxiv
Top 0.1%
5.2%
Show abstract

This study documents the derivation and validation of a deterministic algorithm for cause-of-death (COD) ascertainment from longitudinal real-world medical claims data, evaluated against an independent state-level death certificate file. Death certificates are the dominant reference standard in mortality research but carry well-documented limitations, including primary-cause error rates estimated at 20-40\% across empirical studies. A matched analytic cohort of 216,382 individuals (Connecticut death records, 2017--2025, age 25 and above) was constructed after exclusion of mechanism-of-injury cases and removal of ill-defined symptom-code entries from both sources. Concordance between algorithmic and certificate-based COD was assessed through three complementary frameworks: age-stratified positive predictive value (PPV) at the ICD-10-CM chapter level under a full-set concordance scenario; mean absolute rank difference (MARD) for chapters identified by both sources; and analyses of breadth, depth, and code-level specificity of COD reporting. Chapter-level PPV was strongest for individuals aged 55 and above, with all estimates representing conservative lower bounds given the known error rate of the certificate reference standard. The algorithm consistently reported broader and more granular contributing cause profiles than the death certificate, with discordances directionally consistent with the well-documented tendency of certificates to under-report contributing conditions. These findings support the conclusion that algorithmic COD ascertainment from longitudinal claims data is a feasible and scalable alternative to certificate-based attribution and, at population scale, a principled methodology for characterising death certificate error rates beyond what small-sample chart review studies can achieve.

8
Tacrolimus variability and creatinine predict readmission after liver transplantation

Korenblat, K. M.

2026-07-06 transplantation 10.64898/2026.07.02.26357106 medRxiv
Top 0.1%
4.9%
Show abstract

Unplanned readmissions after liver transplantation occur in over 30% of recipients, yet no validated prediction models exist, and prior observational studies suffer from immortal time bias. The optimal readmission window for outcome prediction and the feasibility of early risk stratification remain undefined. This study is a retrospective analysis of 922 adult liver transplant recipients (August 2018-August 2025) at a single center. Time-varying Cox regression evaluated 14-, 30-, and 90-day readmission windows as predictors of 1-year mortality, correcting for immortal time bias. Gradient-boosted machine learning models leveraging 528,400 laboratory measurements (28 analytes) predicted 90-day readmission using either complete hospitalization data or data restricted to postoperative day 7. Feature importance was quantified by gain, and clinical utility was assessed through risk stratification. Among 902 hospital survivors, 342 (37.9%) experienced an unplanned readmission within 90 days of initial discharge. Only the 90-day readmission window predicted 1-year mortality in time-varying analysis (HR 1.73, 95% CI 1.17-2.57, p=0.006). The model for readmission using complete data achieved AUC 0.614 (95% CI 0.576-0.652); the postoperative day 7 restricted model achieved AUC 0.615 (95% CI 0.577-0.652), with no meaningful performance difference. The tacrolimus coefficient of variation x peak creatinine interaction was the dominant predictor in both the complete model (17.3% importance, rank 1) and the day 7 restricted model (20.4% importance, rank 2). This interaction stratified patients into high-risk (tacrolimus CV >0.3 and creatinine >2.0 mg/dL; 49.8% readmission) versus low-risk (24.8% readmission) groups (risk ratio 2.01, p<0.001). These results identify a modifiable biological determinant of readmission and establish a framework for targeted interventions to reduce unplanned readmission and improve post-transplant outcomes.

9
Projected Population-Level Impact of Digital Return of Results for Cardiovascular-Kidney-Metabolic Screening at US Blood Donation Centers: A Monte Carlo Simulation Study

Qian, Z.; Khera, A.; Makhnoon, S.; Chapman, B. E.; Bryant, B.; Sayers, M.; Compton, F.; Eason, S.; Xing, C.; Ahmad, Z.

2026-09-03 public and global health 10.64898/2026.09.01.26360806 medRxiv
Top 0.1%
4.2%
Show abstract

Background. Cardiovascular-kidney-metabolic (CKM) syndrome affects nearly 90% of US adults, yet most individuals at early, modifiable stages remain unidentified outside clinical care. Blood donation centers offer a scalable, non-clinical venue for CKM screening, but the potential benefit of screening in this context remains unclear. We projected the population-level impact of effective digital return of results (ROR) to inform the design of a pragmatic trial. Methods. We developed a Monte Carlo simulation (100,000 iterations) of the incident major adverse cardiovascular events (MACE), end-stage renal disease (ESRD), and type 2 diabetes (T2DM) preventable by ROR-prompted, guideline-concordant follow-up among donors in CKM Stages 1-2. The estimand counts only events averted by donors who act because of ROR; the intervention effect was modeled directly on strictly positive support, and action was translated into prevented events through a hazard-based cumulative-incidence difference that counts each donor at most once. We evaluated 18 design cells (donor volumes 300,000, 1 million, and 8 million/year; 5- and 10-year horizons; action-rate gains of +10, +20, and +30 percentage points [pp]) and, in a complementary two-arm simulation, the assurance (expected power) of detecting the effect in a single deployment. Results. Under the primary +20 pp scenario, ROR at a single large blood center (300,000 donors/year) is projected to prevent a median of 2,201 events (95% uncertainty interval [UI], 1,099-4,364) over 10 years, scaling to 58,526 (29,154-116,769) at the national donor pool. All 18 design cells had strictly positive 95% lower bounds. The number needed to screen was 136 and the screening cost $2,045 per event prevented (at $15/donor), both invariant to donor volume. Impact scaled linearly with volume and effect size but sub-linearly with the horizon. Detection of the effect was effectively certain at gains of +20 pp or larger (assurance [&ge;]99.6% in every cell and >99.9% in all but the smallest 5-year cell). Conclusions. Even under the conservative scenario, digital CKM ROR at blood donation centers is projected to prevent hundreds to tens of thousands of incident cardiometabolic events at a screening cost per event well within accepted prevention benchmarks, providing prospective, quantitative justification for a pragmatic, randomized evaluation of digital ROR in non-clinical screening settings.

10
Incidence and Management of Early Breakthrough HSV Infection Among Allogeneic Hematopoietic Stem Cell Transplant Recipients

Fischer, M. D.; Mohan, R.; Wald, A. D.; Phipps, A. I.; Ford, E.; Gooley, T.; Tverdek, F.; Biernacki, M. A.; McCulloch, D. J.; Boeckh, M. J.; Johnston, C.; Pergam, S.

2026-08-04 infectious diseases 10.64898/2026.08.02.26359522 medRxiv
Top 0.1%
4.1%
Show abstract

Background: Reactivation of herpes simplex viruses (HSV) can occur in the early post-allogeneic hematopoietic cell transplant (aHCT) period despite antiviral prophylaxis. Few studies have assessed HSV infection in the modern era, in which acyclovir/valacyclovir is recommended for up to 1 year post aHCT. We evaluated the incidence and management of breakthrough HSV during the first 100 days post-aHCT over two decades. Methods: Patients who received their first aHCT at Fred Hutchinson Cancer Center between 2002-2022 were reviewed for breakthrough HSV infection within the first 100 days on prophylaxis (acyclovir 800 mg or valacyclovir 500 mg twice daily). Cases were identified via culture, polymerase chain reaction, and/or direct fluorescent antibody testing; clinical records were reviewed for symptoms, outcomes, and prophylaxis/treatment regimens. Refractory/resistant (R/R) infections were defined according to consensus guidelines. Results: We reviewed data from 4,357 aHCT recipients aged [&ge;]18 years, among whom 3,749 (86%) were HSV seropositive and 23 developed breakthrough HSV infection (observed probability = 0.6%). Among those who had an infection, the median time from transplant to first positive test was 46 days (IQR: 24.0-69.5). Oral and genital mucosa were the most common sites of infection. In total, 11 of 23 (47.8%) patients with breakthrough HSV developed R/R infection. Conclusions: Breakthrough HSV infections are rare in the first 100 days after aHCT among patients receiving antiviral prophylaxis. Refractory/resistant infections were uncommon but represented almost half of breakthrough cases. Our findings highlight the sustained effectiveness of universal prophylaxis in the early post-transplant period.

11
Developing an OMOP-Standardized Prostate Cancer Database and Improving Data Quality Using NLP and PSA-Based Algorithms

Wang, J.; Jackson, J. C.; Garza, A.; Nalla, S.; Ninnemann, T.; Zhang, Y.; Kuo, Y.-F.

2026-07-02 health informatics 10.64898/2026.06.30.26356984 medRxiv
Top 0.1%
4.0%
Show abstract

Objective: To develop and evaluate an Observational Medical Outcomes Partnership (OMOP) standardized prostate cancer database from the University of Texas Medical Branch (UTMB) Epic Electronic Health Record (EHR) and improve data quality using natural language processing (NLP) and prostate-specific antigen (PSA) based algorithms. Materials and Methods: We built a data pipeline to transform UTMB Epic EHR data from 2010 to 2021 into OMOP Common Data Model (CDM) v5.4. Data quality was assessed by comparing the OMOP-standardized data with Galveston Cancer Registry data using availability agreement, Cohen's kappa, and Intraclass Correlation Coefficient. NLP was used to extract PSA, Gleason score, and cancer stage from clinical text, and PSA-based algorithms were used to identify missing treatment and biochemical recurrence. Results: We extracted 815 analytic cases from UTMB EHR. Among them, 700, or 85.9%, were complete and concordant with the cancer registry. PSA showed excellent value agreement. Structured Gleason score and stage data were sparse, with fewer than 20 cases, but NLP greatly improved capture. Treatment agreement was good compared with the cancer registry and improved slightly for radical prostatectomy after applying a PSA-based algorithm. Using PSA trajectories, we identified 60 cases of biochemical recurrence. Discussion: The OMOP-standardized data from UTMB showed good agreement with the cancer registry. However, structured EHR fields incompletely captured diagnosis, pathology, and treatment details. NLP and PSA-based algorithms substantially improved data capture. Manual review also revealed errors in registry data, showing that OMOP-standardized EHR data can complement and help improve cancer registry quality. Conclusion: OMOP standardization combined with NLP and PSA-based algorithms improved prostate cancer data quality and research readiness.

12
Evaluative Stance Toward Artificial Intelligence in High-Quartile Medical Journals (2021-2026): Large-Scale LLM-Assisted Computational Content Analysis

Wang, L.; Poenaru, D.

2026-07-27 health informatics 10.64898/2026.07.23.26358815 medRxiv
Top 0.1%
3.3%
Show abstract

Background: Medical-AI publications do more than report technical performance; they also frame AI as beneficial, uncertain, or risky. How this evaluative stance has changed across the medical literature is not well characterized. Objective: To characterize evaluative stance in published medical-AI discourse abstracts from January 2021 through April 2026 and examine variation over time, concern themes, failure mechanisms, specialties, first-author geography, and publication format. Methods: We conducted an LLM-assisted computational content analysis of medical-AI abstracts from first- and second-quartile medical journals. Of 97,492 post-cutoff Q1/Q2 records entering the prefilter, 16,759 were retained as discourse or evaluative. Claude Sonnet 4.6 assigned 16,749 valid stance classifications using Alarm, Caution, Neutral, Cautious Optimism, and Advocacy. Annual analyses used 16,747 records dated 2021-2026. Critical stance was defined as Alarm plus Caution and indexed evaluative scrutiny rather than opposition or author psychology. Each LLM step was validated against blinded human coding by one author: prefilter Cohen kappa = 0.51, stance quadratic-weighted kappa = 0.79 (95% CI 0.72-0.84) for codable, in-scope records, specialty kappa = 0.75, and mechanism-axis kappa = 0.84 for model type and 0.57 for failure mode. Results: Advocacy declined from 2.9% in 2021 to 0.6% in partial 2026, while Cautious Optimism remained the majority stance. Among 16,749 valid classifications, 30.8% were critical. Critical share increased from 25.4% to 32.6%, a 7.25-percentage-point increase based on unrounded estimates. Among critical records, patient safety remained the most prevalent concern. Hallucination/errors increased by 30.9 percentage points. Regulation declined by 22.0 percentage points and ethics/bias by 8.1 percentage points in prevalence share; these declines do not necessarily indicate lower publication counts. Within the hallucination/error theme, factual error was more common than fabrication. Fabrication estimates should be treated as an upper bound because failure-mode agreement was moderate. Specialty patterns were heterogeneous. Critical rate was inversely associated with FDA-cleared device availability (Spearman rho = -0.65, two-sided p = 0.004), which does not measure adoption, deployment, maturity, or clinical use. First-author geography described publication metadata and discourse, not national attitudes or research quality. Reviews were the least critical and most favourable format. In exploratory forward validation, 2 of 78 early Advocacy predictions were fully borne out, although the analysis was single-rater and retrieval-dependent. Conclusions: Published medical-AI abstracts became modestly less promotional and more focused on specific errors and safety concerns. Unqualified promotion declined, but qualified favourable framing remained dominant, and the rise in critical stance was modest. Concern moved toward errors and patient safety, with factual error discussed more often than fabrication. These findings describe published discourse, not AI capability or whether the evaluations were correct.

13
Rationale and guidance for implementing the continual reassessment method for dose-finding in controlled human infection model studies

Weerasinghe, C.; Osowicki, J.; Simpson, J. A.; Crocker-Buque, T.; McCarthy, J.; Williams, E.; Price, D. J.

2026-07-17 infectious diseases 10.64898/2026.07.16.26358128 medRxiv
Top 0.1%
3.3%
Show abstract

Controlled human infection models (CHIMs) are increasingly used in infectious disease research to study pathogen dynamics and evaluate interventions under controlled conditions. However, these studies are resource-intensive and involve ethical and safety constraints, making efficient study design critical. Dose-finding is a key early component in CHIMs, where the aim is to identify a challenge dose that achieves a target infection probability. Traditional rule-based designs are commonly used but can be inefficient, motivating the use of model-based adaptive approaches such as the Bayesian Continual Reassessment Method (CRM). Although CRM has been extensively studied and widely adopted in Phase I oncology trials for identifying the maximum tolerated dose of therapeutics, its application in CHIM settings remains limited, particularly when the endpoint of interest is infection. This tutorial provides step-by-step guidance for implementing a Bayesian CRM in dose-finding CHIMs, using an oropharyngeal Neisseria gonorrhoeae challenge as a motivating case study. The framework outlines key design components, including dose-grid specification, dose-response model, prior elicitation, Bayesian updating, decision rules, and stopping criteria, with particular emphasis on a clinically interpretable parameterisation. Trial operating characteristics are evaluated through simulation studies under multiple dose-response scenarios and prior-predictive analyses, and compared with a commonly used '3+3' type rule-based design. This work highlights the advantages of Bayesian model-based designs for dose-finding in CHIMs over classic rule-based designs and provides a structured, reproducible framework for implementing CRM, supporting their application in future CHIM studies.

14
County Year Informatics Model for Annual and Cumulative Unique Lung Cancer Screening Eligibility in Maryland, 2026 to 2045

Adebamowo, C.; Adebamowo, S. N.

2026-06-17 epidemiology 10.64898/2026.06.15.26355716 medRxiv
Top 0.1%
2.7%
Show abstract

Purpose: Population-level lung cancer screening programs require denominators that reflect age, smoking history, geography, and changing eligibility over time. We estimated annual prevalent and 20-year cumulative unique low-dose computed tomography screening eligibility for Maryland residents under alternative screening criteria. Methods: We built a deterministic cohort-cell stock-flow simulation using Maryland county-equivalent jurisdiction projections by age, sex, and race/ethnicity, with ACS socioeconomic/nativity covariates and smoking-history priors for ever-smoked status, pack-years, and quit-years. Scenarios included USPSTF 2013 legacy, USPSTF 2021, ACS 2023/2024, a risk-model-expanded sensitivity, and ever-smoked-only capacity stress tests. Cumulative unique eligibility counted people once at first eligibility rather than summing annual prevalent person-years. Results: Under USPSTF 2021, an estimated 238,346 Maryland residents were eligible in 2026 and 245,326 in 2045. The 20-year cumulative unique denominator was 768,668, whereas naively summing annual prevalent counts produced 4,850,735 person-years, a 6.31-fold overcount. ACS 2023/2024 expanded annual eligibility to 314,616 in 2026 and cumulative unique eligibility to 902,796 by adding remote former smokers. Ever-smoked-only adult eligibility was 1,957,699 in 2026 and 3,383,683 cumulative unique over 20 years. Conclusion: A Maryland statewide screening initiative should plan from cumulative unique eligibility and county-equivalent jurisdiction-specific burden rather than annual prevalence alone. Explicit pack-year and quit-year modeling materially changes statewide and county allocation compared with current-smoking proxy models.

15
Interpretable Machine Learning to Improve Donor-Recipient Matching at Time of Heart Transplantation

Xu, J.; Dai, W.; Goldberg, J.; Hu, I.; Chen, C.-H.; Shah, P.; DeFilippi, C.; Sun, J.

2026-07-31 transplantation 10.64898/2026.07.29.26359283 medRxiv
Top 0.1%
2.7%
Show abstract

BACKGROUND: Machine learning (ML) models have been used to evaluate one-year post-transplant mortality in donor-recipient pairs. Previous modeling utilizing noninterpretable ML methods (deep neural networks [DNN] and XGBoost) showed modest gains in area under the receiver operating curve (AUC) beyond logistic regression, but suffered a significant drop in predictive AUC applied to the subsequent years' data and lacked statistical significance in validating previously identified risk factors. METHODS: Using balanced SRTR datasets, we evaluated non-interpretable models against interpretable ML models (adaptive logistic regression with interaction terms [aLR], Classification and Regression tree [CART], and conditional inference tree [CIT]) for one-year mortality, including comprehensive clinician-supervised data curation and inclusion of variables describing pre- and post-2018 listing status changes. Models were trained/tested using rolling-window validation across years and further analyzed with repeated ten-fold crossvalidation. Interaction terms were obtained via Adaptive Best-Subset Selection (ABESS). RESULTS: Predictive validation before the listing policy change in 2018 showed similar AUCs between DNN (0.579), aLR (0.642), CART (0.579), and CIT (0.584), with XGBoost having a higher (0.763) AUC. However, in the post-2018 predictive analysis, aLR outperformed XGBoost (AUC 0.613 vs. 0.586). The interpretable ML models confirm the significance of previously reported risk factors (recipient bilirubin and creatinine) and identify risk factors not previously reported (donor pH, potential recipient distance, and recipient transfusion), as well as clinically relevant interaction terms. CONCLUSIONS: Carefully developed interpretable ML models of one-year transplant mortality have similar predictive performance to black-box models, while identifying novel risk factors, and showing improved performance after recent listing policy changes. With appropriate validation and additional data, interpretable ML modeling may allow real-time data-driven donor selection.

16
Single-Axis Fairness Interventions Produce Asymmetric Cross-Axis Effects in Clinical Prediction

Yoon, K. Y.; Kwak, H.

2026-06-26 health informatics 10.64898/2026.06.16.26355120 medRxiv
Top 0.1%
2.6%
Show abstract

Objective: To systematically evaluate whether single-axis demographic balancing introduces cross-axis fairness trade-offs in clinical prediction, and to characterize the directional asymmetry of these effects across model architectures and balancing strategies. Materials and Methods: We evaluated cross-axis fairness effects of demographic balancing in one-year all-cause mortality prediction using the MIMIC-IV database (N=64,427). Seven machine learning architectures were trained under six balancing strategies targeting either gender or race, with performance assessed via outcome-stratified 80-20 train-test splits repeated 30 times across both targeted and non-targeted axes using AUC, TPR, and Brier score. Results: Gender-targeting interventions largely preserved race fairness, while race-targeting consistently disrupted gender fairness across all methods and the majority of architectures. This asymmetry was invisible to same-axis evaluation alone. Race-targeting also incurred greater performance costs and calibration loss, with observed fairness gains potentially reflecting leveling down rather than genuine improvement. The same intervention could appear successful under TPR but fail under AUC evaluation. Discussion: The asymmetry likely reflects differential category complexity: binary gender balancing requires modest distributional shifts, whereas multi-category race balancing necessitates aggressive reweighting that propagates to correlated axes. Cross-axis fairness effects are directionally dependent and metric-sensitive, indicating that single-metric, single-axis evaluation is insufficient. Conclusion: Single-axis fairness optimization cannot guarantee cross-dimensional equity. Cross-axis, multi-metric fairness evaluation should be integrated into pre-deployment auditing of healthcare artificial intelligence (AI) models.

17
Beyond the Target Product Profile: Design Requirements for Point-of-Need HCV RNA Testing in Harm Reduction Settings

Mata-Robles, S.; Khalaf, K.; Kelley, J.; Chauhan, A.; Balian, L.; Linnes, J. C.; Rodriguez, N. M.

2026-08-17 infectious diseases 10.64898/2026.08.14.26357799 medRxiv
Top 0.1%
2.5%
Show abstract

Point-of-care Hepatitis C Virus (HCV) RNA assays reduce diagnostic turnaround time but depend on benchtop instrumentation and continuous electricity, limiting their deployment in the harm reduction and community settings where confirmatory testing is most needed, as people who use drugs (PWUD) carry a disproportionate share of the HCV burden in the United States. This is a systemic problem in diagnostic development, where decision-making and design requirements overlook point-of-use stakeholders. Closing the gap requires integrating real-world constraints throughout design rather than validating against user needs once a product already exists. Here, we apply a human-centered design (HCD) approach to inform rigorous, stakeholder-derived design requirements, implementation considerations, and early value proposition for a novel point-of-need HCV RNA test intended for deployment in harm reduction and community settings in Indiana. To determine design specifications grounded in real-world context, our objectives were (1) identifying and characterizing context-specific experiences and barriers to HCV testing among higher-risk populations; (2) assessing the perceived benefits and acceptability of the proposed test within real-world settings across direct and indirect user groups; and (3) translating the user needs and contextual constraints into design requirements and implementation considerations that support the test's clinical, operational, and user-centered value. We conducted 18 semi-structured interviews with frontline staff and HCV testing/treatment pipeline experts (n=11) and people who get tested (n=7) across harm reduction organizations, syringe service programs, and community testing settings, analyzed using Rapid Qualitative Analysis guided by the PARRQA framework. Stakeholders responded positively to a single-encounter point-of-need RNA test, and implementation considerations, including funding restrictions, staffing structures, and diverse deployment settings, directly shaped design requirements spanning turnaround time, sample type and volume, portability, result output, target operator, and ease of use. Benchmarking these stakeholder-derived specifications against the FIND Dx HCV target product profile (TPP) showed that stakeholder input confirmed, modified, or extended several TPP criteria and introduced requirements the TPP does not address. Together, these objectives constitute an upstream, evidence-driven design process that translates contextual and stakeholder knowledge into actionable engineering requirements, highlighting the need for diverse stakeholder engagement at all stages of the design process for closing the translation gap between laboratory-validated diagnostic tools and effective point-of-need deployment.

18
Potential Cost of Multi-Cancer Detection Tests to Medicare

Scherer, L. D.; Matlock, D. D.; Cronin, J.; Gritz, M.

2026-08-27 health economics 10.64898/2026.08.24.26361257 medRxiv
Top 0.1%
2.5%
Show abstract

Multi-Cancer Detection (MCD) tests can detect more than 50 different types of cancer using a blood test. Recently passed law in the U.S. guarantees that Medicare will pay for these tests when they are FDA approved and show evidence for clinical benefit. This manuscript provides estimates of the cost of MCD tests to Medicare under different assumptions of cost per test, eligibility, and screening uptake in the eligible population. This manuscript additionally estimates the cost of follow-up testing resulting from false positive results, which are considered avoidable costs caused by the screening test.

19
Timing of S. aureus-related mortality in a large randomized clinical trial: Implications for future study design

Lee, T. C.; Butler-Laporte, G.; Cheng, M. P.; Mertz, D.; Somayaji, R.; Afra, K.; Bai, A.; Chagla, Z.; Daneman, N.; Grant, J. M.; Johnstone, J.; Kandel, C.; MacFadden, D.; Poulin, S.; Prosty, C.; Schwartz, K.; Silverman, M.; Smith, S.; Wuerz, T.; Tong, S. Y.; McDonald, E. G.

2026-06-23 infectious diseases 10.64898/2026.06.20.26356148 medRxiv
Top 0.1%
2.5%
Show abstract

Background: Longer follow-up periods in clinical trials for S. aureus bacteremia (SAB) may capture unrelated deaths, adding random noise that risks biasing trial results towards the null. Objective: To evaluate the timing and infection-relatedness of deaths within a large SAB clinical trial platform. Design: Blinded duplicate adjudication of trial deaths using a modified 7-point Likert-Scale. A third reviewer settled disagreements. Setting: 37 Canadian hospitals participating in the S. aureus Network Adaptive Platform (SNAP) Trial. Participants: 1515 adult patients recruited to SNAP between February 2022 and May 2026. Measurements: Timing and relatedness of 90-day deaths categorized as at least possibly SAB-related not likely to be SAB-related. Optimal follow-up cut-off was determined using Youden's index and graphically. Results: 247 deaths occurred; 97 (39.3%) were adjudicated as at least possibly SAB-related and 150 (60.7%) as not likely related. For probably/definitely related deaths, interrater agreement was 85.0% (Gwet's AC 0.73, substantial); for at least possibly related, it was 77.3% (Gwet's AC 0.55, moderate). Median survival was significantly shorter for SAB-related deaths (12 vs. 30.5 days; difference: 19 days earlier, 95% CI: 12-26, p<0.0001). Nearly 80% of SAB-related deaths occurred by day 30, whereas 50% of unrelated deaths occurred between days 30 and 90. Youden's index optimized follow-up at 20.5 days. Limitations: Potential for cause of death misclassification and data limited to Canadian sites. Conclusion: Deaths considered attributable to SAB cluster rapidly within the first month, while later deaths are predominantly unrelated. A 30-day all-cause mortality window may be more appropriate than 90 days for primary mortality outcomes in trials evaluating acute SAB therapies with longer follow up reserved for metastatic infection and recurrence.

20
CAUSAL-RSV: Causal Analysis of RSV Vaccine Effects in Infants Using Real-World Data

Regan, A. K.; Coates, M. M.; Sullivan, S. G.; Munoz, F. M.; Rowe, S. L.; Avila, C.; Arah, O. A.

2026-07-18 infectious diseases 10.64898/2026.07.16.26356876 medRxiv
Top 0.1%
2.4%
Show abstract

Respiratory syncytial virus (RSV) contributes to substantial morbidity and mortality in young children each year. In 2023, two new prevention products were licensed and recommended in the United States (US), including a prefusion F protein subunit vaccine (RSVpreF) administered during pregnancy and a long-acting monoclonal antibody (mAb) administered in infants. Although post-licensure real-world studies support the effectiveness of RSVpreF vaccine during pregnancy, existing studies have been conducted in settings where only RSVpreF vaccine is available. The real-world effectiveness of RSVpreF vaccine in settings where both RSVpreF vaccine and mAbs are available is not yet well understood. The goal of this study is to estimate the real-world effectiveness of the RSVpreF vaccine against severe infant RSV by applying causal mediation analysis with receipt of mAbs as a mediating variable. Using a national cohort of mother-infant dyads with the Optum Labs Data Warehouse (OLDW), we will model vaccine and mAb effects in a longitudinal cohort spanning the 2023-24, 2024-25, and 2025-26 RSV seasons. Results will be used to better understand the total effect of RSVpreF vaccination when it is used as one component within a hybrid infant RSV prevention program.